Papers with SOTA zero-shot image captioning

1 papers
Text-Only Training for Image Captioning using Noise-Injected CLIP (2022.findings-emnlp)

Copied to clipboard

Challenge: a new approach to image captioning requires large datasets of captioned images and is difficult to collect.
Approach: They propose to use a decoder to translate CLIP textual embeddings back into text . they show that this intuition is “almost correct” because of a gap between the embeddable spaces .
Outcome: The proposed approach shows that the intuition is “almost correct” because of a gap between the embedding spaces, and rectifies this via noise injection during training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations